Journal of Clinical Epidemiology
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Journal of Clinical Epidemiology's content profile, based on 31 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.
Show abstract
Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.
Oparah, C.; O'Keefe, H.; Agbeleye, O.; Nesworthy, J.; Norman, G.; Kunonga, T. P.
Show abstract
Clinical trials often enrol populations that differ from those who ultimately receive the interventions, raising concerns about external validity and health equity. Trial registries could provide an early opportunity to assess representativeness, but it is unclear whether registry data contain sufficient information to enable such assessments. This study evaluated the feasibility of using registry data to assess representativeness in Phase II and III pharmacological randomised controlled trials. A search of ClinicalTrials.gov from December 2024 to January 2025 identified trials with results posted after 1 January 2023 across cardiovascular disease (CVD) excluding stroke, diabetes mellitus, and selected mental health disorders. Of 1,328 records screened, 98 trials met inclusion criteria (51 Phase III, 47 Phase II). Reporting completeness was variable, particularly in Phase II studies. CVD and diabetes trials predominantly included middle-aged to older adults, while mental health trials recruited mainly individuals aged 36 to 50 years. Across CVD and mental health trials, participants were largely male. Reporting of BMI, contraception, and comorbidity criteria was inconsistent, though available data suggested these factors influenced sample composition. Fewer than 10% of trials reported equity-relevant characteristics beyond age and sex, and none addressed intersectionality. Assessing equity using registry data is feasible but constrained by incomplete and inconsistent reporting.
Cummins, J.; Drysdale, H.; Elson, M.; Hussey, I.; Goldacre, B.; DeVito, N. J.
Show abstract
Objective To evaluate the accuracy and cost of RegCheck, an automated large language model (LLM)-based workflow, for identifying clinical trial outcomes and detecting outcome misreporting by comparing its outputs with manual assessments from the COMPare Trials project. Design Validation study. Setting Sixty-two clinical trials originally assessed in the COMPare Trials project, sampled from five high impact general medical journals. Participants Published clinical trial reports and their corresponding prespecified registrations and/or protocols. Main outcome measures Four prespecified research questions were examined. RQ1 assessed outcome extraction recall relative to COMPare. RQ2 assessed accuracy of outcome classification as primary, secondary, or non-prespecified. RQ3 assessed accuracy of misreporting detection relative to COMPare, with additional manual adjudication of discrepancies between RegCheck and COMPare. RQ4 assessed the average per-paper cost of running the automated workflow. Results Across the validation papers, RegCheck achieved 91.2% outcome extraction recall relative to COMPare, and 83.6% outcome classification accuracy. For detection of outcome misreporting, RegCheck's overall accuracy was 85.6%. However, after resolving discrepancies with the original human judgements (which frequently favoured RegCheck's judgement), revised accuracy for outcome misreporting detection was 94.8%. The mean cost of running the workflow was 5.94 USD per paper. Conclusions RegCheck achieved high overall performance with a rigorous manual benchmark for identifying prespecified and reported outcomes in clinical trials, and detecting outcome misreporting, while operating at very low marginal cost. Adjudication of discrepant judgements suggested that RegCheck frequently identified valid issues not captured in the reference standard. Automated outcome checking may offer a scalable way to support editors, peer reviewers, and authors in detecting outcome switching and improving trial reporting.
Jajieh, T.; Chapelle, C.; Lemarchand, C.; Locher, C.; Pencole, M. A.; Ioannidis, J.; Ollier, E.; Naudet, F.; Laporte, S.
Show abstract
Background: Randomized controlled trials (RCTs) that rely on fabricated or fictitious data, commonly known as zombie trials, threaten evidence-based medicine by contaminating systematic reviews and clinical guidelines. We describe a cohort of zombie trials with confirmed data fabrication. Methods: The cohort was derived from the Retraction Watch database by updating and extending the dataset used in the VITALITY study to include all retracted zombie trials identified from database inception through December 11, 2025. To ensure that included RCTs were identified as zombie trials rather than retracted for other reasons, retraction notices were text-mined for terms indicative of data fabrication. Trial- and author-level characteristics were extracted using a combination of automated and manual methods by one reviewer. We applied descriptive statistics, coauthorship network analysis, and cascade effect mapping. Results: Among 236 retracted RCTs published between 1983 and 2024 and identified as zombie trials with confirmed data fabrication, 197 (83.5%) were associated with authors who had a history of repeated misconduct. These trials involved 450 authors overall, including only 66 unique first authors, with one author alone accounting for 104 trials (43.8% of the cohort). Authors associated with multiple zombie trials experienced substantially longer delays to retraction, with a median of 13.6 years (IQR 8.0-15.3), compared with 2.75 years for authors linked to a single zombie trial. Most retracted zombie RCTs were monocentric (91.9%), involved a median of 3.5 authors per trial, and predominantly evaluated pharmacological interventions. They were largely concentrated in anesthesiology (58.1%) and in Japan (56.4%). Adjusted for the volume of trials produced by country, Japan, Tunisia, and Egypt had the highest rates. Coauthorship networks formed fragmented, largely disconnected clusters, suggesting localized factories of fabricated evidence. Retraction cascades were common, with the identification of a single fraudulent trial often leading to chains of retractions that exposed misconduct spanning decades. Conclusion: Confirmed zombie trials were usually generated from single centers that published many such trials.
Zhang, L.; Zhang, W.; Li, L.; Ye, Y.; Zhu, X.; Shao, L.; Yang, H.; Hu, Y.; Li, Y.; Lin, H.; Geng, W.
Show abstract
Abstract Background Persistent pain in later life is associated with functional decline. Identifying resources associated with preserved daily function may broaden pain research beyond symptom burden alone. Objective To examine six prespecified resource indicators-physical activity, depressive symptoms, recall performance, current partner status, education and wealth-in relation to subsequent maintenance of independence in activities of daily living (ADL). Methods This longitudinal observational study analysed adults aged 50 years or older with pain at both the initial and baseline assessments in the Health and Retirement Study (HRS), the English Longitudinal Study of Ageing (ELSA) and the Survey of Health, Ageing and Retirement in Europe (SHARE). ADL maintenance was defined as no difficulty in all five prespecified ADL activities at baseline and no difficulty in all five at follow-up. Indicator-specific risk ratios were estimated using modified Poisson regression with robust variance and pooled using random-effects meta-analysis. The six meta-analysis P values were adjusted using the Holm procedure. Each indicator used a separate complete-case sample; no common six- indicators sample was constructed. Results The descriptive samples comprised 8,279 HRS person-windows contributed by 4,688 unique participants, 1,200 ELSA participants and 6,995 SHARE participants; indicator-specific denominators varied. Physical activity was associated with a higher probability of ADL maintenance across all three cohorts (pooled risk ratio, 1.145; 95% CI, 1.097-1.195; P = .005; Holm-adjusted P = .032) and was the only prespecified hypothesis to meet the Holm-adjusted criterion. The other five hypotheses did not meet this criterion; failure to do so should not be interpreted as evidence that the corresponding associations were absent. Several of these pooled estimates showed substantial heterogeneity. Conclusions Among older adults with persistent pain, physical activity was associated with a higher probability of ADL maintenance, and the physical activity hypothesis was the only one of the six prespecified hypotheses to meet the Holm-adjusted criterion in the three-cohort meta-analysis. These observational findings support further study of physical activity and functional maintenance but do not establish causal benefit.
LIn, H.; Lyu, J.
Show abstract
BackgroundQuality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. ObjectiveTo estimate inter-model reliability for QCC quality scoring and to assess whether model scores align with designed synthetic anchors and remain descriptively comparable to a small set of public PMC QCC reports. MethodsWe evaluated 30 synthetic QCC reports and 8 public PMC QCC reports across four primary evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because it shared the model family used during prompt development. Each synthetic case was scored across eight QCC quality dimensions in three runs per evaluator. We summarized each evaluator by median scores, then estimated ICC(A,1) across the primary panel. We also examined score-based calibration against designed anchors, keyword-assisted defect mention, leave-one-out and k=5 sensitivity, and a descriptive synthetic-versus-PMC distributional plausibility check. ResultsInter-model reliability on the primary k=4 panel was excellent: ICC(A,1) = 0.953 (95% CI 0.944 to 0.962) with 237 pooled case-dimension rows. The pre-specified k=5 sensitivity analysis including Claude was 0.954, and leave-one-out estimates within the primary panel ranged from 0.950 to 0.959. Score-based calibration against designed anchors met the prespecified target in 57/58 trap-affected case-dimension rows (98.3%). Keyword-assisted defect mention was present in 51/58 trap instances (87.9%). The synthetic-versus-PMC comparison was descriptively similar across all eight dimensions, and all dimensions met the predefined descriptive margin check. ConclusionsIn this designed-anchor pilot, multi-model LLM scoring of QCC methodological quality showed high inter-model reliability and stable alignment with synthetic anchor scores. These findings support benchmark feasibility, but they do not establish expert validity, clinical validity, or operational deployment readiness.
Ledley, F. D.; Mozer, R.
Show abstract
There were substantial changes in NIH policies regarding research funding in FY2025. This work examines NIH funding for pediatric research FY2020 through Q2FY2026 including the number and cost of awards, the number of first year (type 1) awards, the number of Notices of Funding Opportunity, and the topical focus of research awards. NIH funding for pediatric research declined >20% in the first two quarters FY2025-FY2026 with proportionally greater reductions in first year awards and Notices of Funding Opportunity. Changes were also noted in the topic prevalence of new awards consistent with 2025 guidelines identifying topics "not aligned with NIH priorities." These results suggest that pediatric research aimed at advancing healthcare for children is at risk with potential collateral consequences beyond childhood.
Cumpston, M. S.; Brennan, S. E.; Ryan, R.; Thomas, J.; McKenzie, J. E.
Show abstract
Introduction Systematic review authors commonly encounter situations where the data required for meta-analysis are incompletely reported (e.g. when effect estimates are reported without a measure of precision). In this circumstance, many systematic review authors use a method other than meta-analysis (e.g. vote counting), but rarely describe those methods or the rationale for selecting them. We aimed to investigate what methods authors consider when meta-analysis of all study results is not possible, and what factors influence their decisions. Methods We interviewed 12 experienced systematic review authors, editors and methodologists, presenting four scenarios in which it was not possible to combine all results using meta-analysis. Scenarios varied in the number and size of included studies, available data, and risk of bias. Participants discussed the methods they considered to summarise, synthesise and present the results; whether they would synthesise available results; and how they would draw overall conclusions. Results Factors that informed decisions included participants' overall purpose in conducting synthesis, existing beliefs about study results and synthesis methods, trust in the available data, and the decision-making needs of end users. Participants differed in which synthesis methods to use, whether they would use multiple synthesis methods, and which studies they would analyse with each method. Conclusions We identified several synthesis methods considered when meta-analysis of all results is not possible, and factors that influence the selection of methods, neither of which are routinely reported. More complete reporting of these methods and the factors informing decisions would allow readers to better understand the decisions made.
Tzimas, G.; Vanghelof, J. C.; Mohammed, A.; Raicu, D. S.; Du, L.; Ernst, M. E.; Warner, E. T.; Chan, A. T.; Ryan, J. C.; Espinoza, S. E.; Murray, A.; Sheets, K.; Tchoua, R. B.; Shah, R. C.
Show abstract
Importance: The ASPREE randomized trial found no overall benefit of low-dose aspirin for disability-free survival among older adults. However, individual estimates in pre-specified subgroups indicated potential benefit among racial and ethnic minoritized participants in the United States (US). Objective: To evaluate whether the effect of low-dose aspirin vs placebo on disability-free survival differed across US Black and Hispanic ASPREE participants using individualized treatment-effect estimation. Design, Setting, and Participants: Post hoc clinical trial analysis of ASPREE, a randomized, double-blind, placebo-controlled clinical trial of daily low-dose aspirin vs placebo. This analysis included US ASPREE participants who self-identified as non-Hispanic Black or Hispanic, were aged 65 years or older, and had complete baseline predictor and outcome data. Interventions: Randomization to daily 100-mg aspirin or placebo. Main Outcomes and Measures: The primary outcome was loss of disability-free survival, defined as death, persistent physical disability, or dementia. Individualized treatment effects were estimated post hoc using a Random Survival Forest X-learner. Heterogeneity was evaluated on the relative scale with Cox proportional hazards models and on the absolute scale with 5-year risk differences. Results: Among 2411 US ASPREE participants, 1270 were included in the Black and Hispanic analytic cohort (897 non-Hispanic Black and 373 Hispanic participants; mean age, 71.8 years). Aspirin was associated with lower risk of disability-free survival loss compared with placebo (hazard ratio [HR], 0.65; 95% CI, 0.45-0.93). In model-derived tertiles, aspirin was associated with lower risk in the greatest predicted-benefit group (HR, 0.36; 95% CI, 0.19-0.71; 5-year absolute risk difference [ARD], -11.1 percentage points; 95% CI, -22.0 to -0.1) but not in the lowest predicted-benefit group (HR, 1.26; 95% CI, 0.70-2.27; ARD, +3.9 percentage points; 95% CI, -5.9 to 13.6). Conclusions and Relevance: In these analyses of US Black and Hispanic ASPREE participants, aspirin effects on disability-free survival appear to be heterogeneous, with benefit concentrated in a subset of participants. Because these findings are from post-hoc models, they should be externally validated before being incorporated into clinical decision-making. Trial Registration: ClinicalTrials.gov Identifier: NCT01038583; https://clinicaltrials.gov/study/NCT01038583
McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.
Show abstract
This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.
Papas, K.; Banerji, A. I.
Show abstract
Background. Hypermobile Ehlers-Danlos syndrome (hEDS), postural orthostatic tachycardia syndrome (POTS) and mast cell activation syndrome (MCAS) are reported to co-occur frequently. It is unclear how strongly, whether symptom profiles sharpen prediction of a second diagnosis given a first, and whether the published literature can support such inference at all. Methods. We pooled 22 published cohorts-aggregate prevalence data, no primary human-subjects data-using Bayesian hierarchical random-effects models on the logit scale, and propagated the resulting posteriors through naive and tempered symptom updating. We introduce a feasibility screen derived from the Frechet-Hoeffding bounds that tests whether separately pooled marginals can describe a single population, and we characterise the identifiability of latent class structure under disease-selected sampling. Results. Directed comorbidity is strongly asymmetric: {pi}POTS|hEDS = 46.6% (95% CrI [32.5,61.5]) against {pi}hEDS|POTS = 12.1% ([3.5,38.5]), a near-fourfold gap, with {pi}MCAS|POTS lowest at 3.9% ([0.7,21.6]). Prediction intervals exceed credible intervals throughout, indicating substantial between-cohort heterogeneity. The feasibility screen finds 26 of 102 testable cells (25.5%) incompatible with any joint distribution; critically, 21 of these fail the upper Frechet bound and are invisible to the one-sided screen that is the natural first implementation. Among cells surviving the screen, symptom evidence is informative in four of six directions-P(hEDS | POTS,S) rises from 12.1% to 49.7% on a four-symptom panel under tempered updating, and P(POTS | MCAS,S) from 49.5% to 82.1%-but inert in both hEDS-cohort directions. A pathway-dispersion contrast excludes zero in two of six directions, in opposite signs and by margins of 0.1-0.2 percentage points, consistent with chance at this number of comparisons. We show latent class structure is not identified from disease-selected aggregate data, and that the single cohort reporting trivariate structure (N = 8) yields an exactly balanced table (OR = 1.00, 95% CI [0.063, 15.99]). Conclusions. The pooled directional probabilities are usable as clinical priors, with intervals wide enough to preclude precision. Symptom-conditioned prediction is supported in some directions but not those most often invoked clinically, and every estimate rests on cohorts dominated by self-reported ascertainment. The principal methodological contribution is the two-sided feasibility screen: applied here it shows that a quarter of the testable literature cannot describe one coherent population, and that a one-sided implementation understates this sixfold. Keywords: hypermobile Ehlers-Danlos syndrome; postural orthostatic tachycardia syndrome; mast cell activation syndrome; comorbidity; Bayesian meta-analysis; random-effects model; Frechet bounds; identifiability; latent class analysis; collider bias
Reddy, S.; Heritier, A.
Show abstract
The rapid expansion of the medical artificial intelligence (AI) literature has outpaced our ability to judge how far published models have progressed towards clinical use. We investigated whether the translational maturity of a study can be estimated automatically from its abstract. Using PubMed, we assembled a corpus of 11,024 candidate articles, reduced it to 1,816 AI-related articles by heuristic filtering, and manually double-annotated a balanced sample of 524 articles across five maturity classes (internal validation, external validation, prospective evaluation, implementation or governance, and not applicable). Abstracts were represented as TF-IDF features and classified using multinomial logistic regression with a Lasso penalty, chosen for interpretability and suitability for a small, imbalanced dataset. On a stratified held-out test set (n = 104), the model achieved 69.2% accuracy, Cohen's kappa of 0.495, macro-F1 of 0.458 and a weighted AUC of 0.820. Performance was strong for the frequent classes but poor for the rare implementation or governance class, which the model failed to recover. A balanced manual verification of 200 large-corpus predictions confirmed this pattern, with per-class precision ranging from 82.5% (internal validation) to 5.0% (implementation or governance). An interpretable, low-resource classifier can support literature mapping but requires human oversight for advanced maturity levels.
Jankovic, D.; Palmer, S.; Callister, M. E. J.; Lyratzopoulos, G.; Dias, S.; Welton, N. J.; Payne, K.; Soares, M. O.
Show abstract
Preclinical cancer sojourn time, defined here as the duration a cancer is undetected but detectable, is important for understanding disease progression and evaluating screening policies. This study aims to robustly characterise empirical evidence and existing knowledge over mean sojourn times across 21 stageable tumour sites, including stage-specific preclinical cancer sojourn times and the sojourn time of circulating tumour DNA (ctDNA)-positive cancers. We updated an existing systematic review through to February 2025 to extract population-level empirical sojourn time estimates derived from mathematical models of primary screening data. To synthesise this heterogeneous literature, quantify uncertainty, and obtain estimates for cancer-sites lacking empirical evidence, we conducted a formal Structured Expert Elicitation involving 15 clinical experts. The elicitation was grounded on the systematic review results, supplemented by an evidence dossier that included survival data and outcomes from relevant ctDNA cancer studies. The systematic review revealed heterogeneity in existing literature, which focused on a small subset of screened cancers (e.g., breast, cervical, colorectal). The elicitation successfully generated comprehensive probability distributions of overall mean sojourn times for all 21 cancer-sites (representing the site of tumour origin), as well as stage-specific sojourn times and overall sojourn times for ctDNA-positive cancers across 14 cancer-sites. This study used robust methodology to quantitatively describe existing evidence and experts' beliefs on the sojourn time of multiple cancer-sites, also describing uncertainty. Such estimates are important for future evaluations of the clinical impact, potential for overdiagnosis and subsequent cost-effectiveness of emerging screening technologies, including multi-cancer detection tests.
Jiang, L.; Ying, X.; Brown, A. W.; Lan, M.; Song, W.; Menke, J.; Vorland, C.; Mayo-Wilson, E.; Kilicoglu, H.
Show abstract
Randomized controlled trials (RCTs) play a central role in assessing the benefits and harms of interventions. Incomplete reporting in RCT publications can compromise the verifiability and usefulness of RCTs. SPIRIT and CONSORT reporting guidelines aim to improve the completeness of RCT protocols and results publications, respectively. However, many RCTs are not reported completely. Checking manuscripts automatically could help authors improve the completeness of reports prior to publication. We previously annotated SPIRIT-CONSORT-TM, a corpus of 200 articles (comprising 100 protocol-results publication pairs) using 83 checklist items drawn from SPIRIT 2013 and CONSORT 2010. We also trained machine learning models to automatically assess reporting at the item level. Each checklist item can include multiple constituent elements (i.e., specific details required for that item), and an item might be considered fully reported when all of its elements are present. However, prior work does not explicitly capture or evaluate reporting at the element level. To address this gap, we extended SPIRIT-CONSORT-TM by incorporating element-level annotations and using them to assess reporting completeness (SPIRIT-CONSORT-ELM). We formulated element-level assessment as a machine reading comprehension task, operationalized through 119 questions, where each question targets a specific reporting element within a checklist item. Using the 200 articles included in SPIRIT-CONSORT-TM, two annotators independently answered 119 questions for 50 articles (25 protocol-results pairs) and resolved any discrepancies through discussion; the remaining 150 articles (75 protocol-results pairs) were assessed by a single annotator. We then developed an automated pipeline for element-level assessment using SPIRIT-CONSORT-ELM. The pipeline first applies a PubMedBERT-based model to identify sentences containing item-level reporting information, then it uses a generative large language model (LLM; GPT-5) with chain-of-thought reasoning to answer element-level questions based on the retrieved evidence. Agreement between the two annotators was high (Gwet's AC1: 0.782) and our pipeline achieved high accuracy in identifying element-level reporting evidence (F1: 0.822, Gwet's AC1: 0.796). Ablation studies indicate that chain-of-thought reasoning and the inclusion of illustrative in-context examples modestly improve LLM performance on the machine reading comprehension task. SPIRIT-CONSORT-ELM provides a benchmark for evaluating reporting guideline completeness at the element level, enabling assessment of RCT transparency beyond the simple presence or absence of checklist items and is publicly available at https://osf.io/kznx4/. The automated pipeline establishes a robust baseline for assessing RCT reporting and demonstrates potential as a practical aid for authors, reviewers, and editors to identify and address gaps in completeness and transparency of RCT reports.
Aborageh, M.; Korcinska Handest, M. R.; Bakos, I.; Rajamaki, B.; Silva, C.; Horvath-Puho, E.; Pylkkaenen, L.; Venda, C.; Lentzen, M.; Becker, C.; Fernandes, J.; Paakinaho, A.; Vo, T.; Haenisch, B.; Hartikainen, S.; Tolppanen, A.-M.; Furtado, C.; Froehlich, H.; Ehrenstein, V.
Show abstract
Background: Real-world data (RWD) from different countries are increasingly used to support regulatory, health technology assessment (HTA), and population-level evidence generation. However, cross-country analyses are challenged by differences in data provenance, healthcare systems, coding practices, completeness, and clinical workflows. The Observational Medical Outcomes Partnership (OMOP) common data model (CDM) is widely used to harmonise heterogeneous RWD sources, but its ability to improve comparability of downstream epidemiological analyses relative to native source data across countries requires empirical evaluation. Methods: We examined RWD from Denmark, Finland and Portugal in their ability to capture epidemiology of female breast cancer (BC) and amyotrophic lateral sclerosis (ALS), exemplifying, respectively, a common disease with established treatment modalities and high survival and a rare fatal disease with scarce treatment options. To enable head-to-head comparison on a semantic level, data were mapped to the OMOP CDM. Data in the native format were used for comparison. In a downstream analysis, we examined disease epidemiology, patient characteristics, treatment, and survival. Results: OMOP conversion enabled a common analytical framework across countries and supported semantically aligned comparisons of key epidemiological and clinical variables. However, cross-country comparability was influenced by differences in data provenance, population coverage, coding practices, availability of clinical details, treatment capture, and healthcare-system-specific workflows. Iterative comparison with native data and external clinical evidence was necessary to identify mapping issues, assess information loss, and ensure high semantic fidelity of the converted data. Overall, OMOP-based estimates were highly consistent with native-data analyses and existing clinical expectations, but residual discrepancies reflected both source-data heterogeneity and decisions in the Extract, Transform, Load (ETL) workflow design. Conclusions: OMOP CDM conversion facilitates semantically meaningful cross-country analyses of RWD by mapping heterogeneous source data to a common structure and standardised vocabularies. However, CDM conversion does not eliminate heterogeneity in the underlying data-generating processes and cannot substitute for study-specific data quality and fitness-for-purpose assessment. Robust use of harmonised RWD for regulatory, HTA, or population-level evidence generation requires iterative benchmarking against native data, clinical expertise, and data-science expertise to support valid interpretation across countries.
MUTHUKA, J. K.; Nyambura, L. W.; Onyango, C. K.; Oluoch, K.; Kioko, M.; Maluki, J.; Nzioki, J. M.; Kim, S.
Show abstract
Background: Autism spectrum disorder (ASD) is a lifelong neurodevelopmental condition for which timely diagnosis is critical to early intervention, family support, and equitable access to care. However, substantial disparities in access to ASD diagnostic services persist across socioeconomic, geographic, clinical, and health-system contexts. This systematic review and meta-analysis synthesized evidence on determinants of access across the ASD diagnostic pathway, from recognition and referral to diagnostic completion and timely diagnosis. Methods: We systematically searched MEDLINE/PubMed, Embase, Scopus, Web of Science, Global Health, and grey-literature sources for studies published between January 2004 and December 2024. Eligible studies examined determinants of ASD diagnostic completion, diagnostic pathways, diagnostic timeliness, or barriers and facilitators to diagnostic access. Two reviewers independently extracted data and assessed methodological quality using the Mixed Methods Appraisal Tool (MMAT). Quantitatively comparable estimates were synthesized using random-effects models with restricted maximum likelihood estimation. Heterogeneity was assessed using Cochran's Q, I2, tau2, and 95% prediction intervals. Pre-specified subgroup analyses, meta-regression, sensitivity analyses, funnel-plot assessments, and Bayesian random-effects analyses were undertaken. Results: The search identified 4,899 records; after removal of 537 records without associated data, 4,362 records underwent title/abstract screening. 3,800 records were excluded, 562 reports were sought for retrieval, and 450 full-text reports were assessed after 112 could not be retrieved. Ultimately, 22 unique studies met the inclusion criteria. Nine unique studies contributed 23 quantitative effect estimates, while the remaining studies contributed to the narrative synthesis. The evidence covered socioeconomic, geographic, family, communication, screening, child developmental, provider, and health-system determinants. The overall random-effects meta-analysis yielded a pooled diagnostic access outcome of 74.1% (95% CI 65.8-81.1%), with substantial heterogeneity (Qe=209.95, p<0.001; I2=88.4%, 95% CI 79.1-94.4%; tau2=0.691) and a wide 95% prediction interval of 32.8-94.4%. Bayesian analysis produced a highly concordant pooled estimate of 73.3% (95% CrI 65.3-80.2%), with I2=87.5% and tau=0.833, and satisfactory MCMC convergence (R-hat=1.000). By outcome domain, pooled successful outcomes were highest for diagnostic pathways (89.3%, 95% CI 70.1-96.7%), followed by timely diagnosis (76.3%, 95% CI 62.9-86.0%), and lowest for diagnostic completion (67.1%, 95% CI 61.8-72.0%) (Qm=5.98, p=0.050). Timely diagnosis demonstrated particularly high heterogeneity (I2=91.2%), whereas diagnostic completion showed moderate heterogeneity (I2=40.6%). Across determinant domains, frequentist pooled estimates were 79.5% for child developmental/neurobehavioral factors, 74.2% for family/socioeconomic/perceptual factors, 68.0% for intervention/care-navigation factors, and 63.6% for provider/clinical recognition factors. Bayesian estimates were 76.7% (BF=53.76), 72.9% (BF=226.32), 64.3% (BF=25.60), and 53.7% (BF=0.684), respectively. Meta-regression indicated that determinant category (Qm=13.48, p=0.004) and effect measure (Qm=7.81, p=0.020) significantly explained between-study variation, whereas age group (p=0.203) and geographic region (p=0.453) did not. Family/socioeconomic factors had significantly larger effect sizes (B=2.703, 95% CI 0.661-4.744; p=0.009), as did child developmental/neurobehavioral factors (B=1.516, 95% CI 0.047-2.985; p=0.043). Potential small-study effects were detected by two of three asymmetry tests, although the Rosenthal fail-safe N was 1,723. Trim-and-fill identified seven potentially missing estimates, with an adjusted pooled effect of 68.4% (95% CI 27.7-109.1%). Importantly, exclusion of two influential outlying estimates produced a pooled outcome of 77.1% (95% CI 71.6-81.9%), indicating that the principal finding was robust. Conclusions: Approximately three-quarters of observed ASD diagnostic outcomes represented successful access, but the substantial heterogeneity indicates that diagnostic access is highly context-dependent. Families were more likely to successfully navigate diagnostic pathways than to complete diagnostic assessment, while timely diagnosis showed the greatest variability across settings. Family and socioeconomic circumstances and child developmental characteristics emerged as particularly important determinants, whereas provider-related effects were more heterogeneous and uncertain. Improving equitable ASD diagnosis requires interventions spanning the entire diagnostic pathway, including developmental surveillance, screening, referral coordination, family navigation, provider capacity, specialist availability, and mechanisms to ensure completion of diagnostic assessment. Greater longitudinal and implementation research is particularly needed in low- and middle-income countries, where diagnostic infrastructure and specialist capacity remain limited.
Weng, Y.; Yalamaddi, H.; Fu, D.; Mishra, A.; Bunning, B. J.; Martin, A. B.; Hope, J.; Charu, V.; Kurian, A.; Desai, M.
Show abstract
Introduction: For oncology patients with limited treatment options, clinical trials may be a critical lifesaving pathway. Identifying relevant trials, however, is a time-consuming and difficult task. Several patient-trial matching processes incorporating large language models (LLMs) have been proposed to alleviate the burden on patients and oncologists. We aim to explore the benefits and practical challenges of zero-shot LLM-assisted trial matching processes by analyzing the results for a single pancreatic cancer patient. Materials and Methods: The results of a simple zero-shot LLM-assisted clinical trial matching process for our patient were compared to those of a "human benchmark," which was developed manually by two of the authors interfacing directly with ClinicalTrials.gov. Performance metrics -- sensitivity, specificity, precision, and accuracy -- were calculated. In addition, a qualitative content analysis (QCA) of LLM reasoning text was done to identify patterns in "errors," which we define as a human-LLM discrepancy in final patient eligibility. Implications and severity of errors are discussed. Results: The zero-shot LLM-assisted process returned potential trials with a sensitivity, specificity, and precision of 81.1%, 89.3%, and 86.5% respectively compared to the human benchmark. Qualitative error analyses revealed that about 73% of errors could potentially be alleviated with improved prompting and information access. Overall performance seemed comparable to that of human reviewers. Conclusion: The results from this preliminary real-world case study provide additional evidence to the literature in support of the integration of LLMs in clinical trial matching to provide benefit to patients with metastatic cancer with limited options.
Samaan, S.; Devi, J.; Vincent, M.; Coombs, S.; Sehgal, P.; Mouhamed, M.; Rai, V.; Johnson, A. M.; Yarur, A. J.; Barnes, E. L.; Deepak, P.
Show abstract
Background: Large language models (LLMs) offer promise for systematic review data extraction, but performance in complex multidisciplinary domains and utility for clinical statement generation remain insufficiently described. Objectives: To evaluate Google NotebookLM for AI-assisted data extraction and RAND/UCLA consensus statement generation in a systematic review of IBD, obesity, and cardiometabolic comorbidities. Methods: Studies were organized into domain-specific notebooks; structured prompts generated standardized evidence tables. Two independent reviewers validated outputs against full-text articles using a four-category error classification. Cell-level accuracy and critical accuracy (cells free of major factual errors) were the primary metrics; workflow time was compared against a published conventional extraction benchmark. Concordance between AI-generated and expert-finalized statements was assessed. Results: Across 57 articles, 1,710 data cells were extracted; 151 (8.83%) were flagged, yielding 91.17% cell-level accuracy. Major factual errors occurred in only 4 cells (0.23%), for a critical accuracy of 99.77%. Most errors were minor omissions (59.6%) or incomplete extractions (30.5%); domain error rates ranged from 7.08% to 11.33%. The pipeline required 17.7 versus a projected 165.1 person-hours (89.3% reduction). PICO-structured prompting generated 70 candidate statements; 58 of 112 finalized panel statements (51.8%) were AI-derived, and 85.7% were retained in the finalized set. Conclusion: Google NotebookLM demonstrates feasibility as a primary extraction and synthesis tool in a multidisciplinary systematic review, with extractive incompleteness as the principal limitation and substantial time savings over conventional approaches. Its novel application to RAND/UCLA consensus statement generation extends AI-assisted evidence synthesis to clinical consensus generation workflow.
Wang, L.; Poenaru, D.
Show abstract
Background: Medical-AI publications do more than report technical performance; they also frame AI as beneficial, uncertain, or risky. How this evaluative stance has changed across the medical literature is not well characterized. Objective: To characterize evaluative stance in published medical-AI discourse abstracts from January 2021 through April 2026 and examine variation over time, concern themes, failure mechanisms, specialties, first-author geography, and publication format. Methods: We conducted an LLM-assisted computational content analysis of medical-AI abstracts from first- and second-quartile medical journals. Of 97,492 post-cutoff Q1/Q2 records entering the prefilter, 16,759 were retained as discourse or evaluative. Claude Sonnet 4.6 assigned 16,749 valid stance classifications using Alarm, Caution, Neutral, Cautious Optimism, and Advocacy. Annual analyses used 16,747 records dated 2021-2026. Critical stance was defined as Alarm plus Caution and indexed evaluative scrutiny rather than opposition or author psychology. Each LLM step was validated against blinded human coding by one author: prefilter Cohen kappa = 0.51, stance quadratic-weighted kappa = 0.79 (95% CI 0.72-0.84) for codable, in-scope records, specialty kappa = 0.75, and mechanism-axis kappa = 0.84 for model type and 0.57 for failure mode. Results: Advocacy declined from 2.9% in 2021 to 0.6% in partial 2026, while Cautious Optimism remained the majority stance. Among 16,749 valid classifications, 30.8% were critical. Critical share increased from 25.4% to 32.6%, a 7.25-percentage-point increase based on unrounded estimates. Among critical records, patient safety remained the most prevalent concern. Hallucination/errors increased by 30.9 percentage points. Regulation declined by 22.0 percentage points and ethics/bias by 8.1 percentage points in prevalence share; these declines do not necessarily indicate lower publication counts. Within the hallucination/error theme, factual error was more common than fabrication. Fabrication estimates should be treated as an upper bound because failure-mode agreement was moderate. Specialty patterns were heterogeneous. Critical rate was inversely associated with FDA-cleared device availability (Spearman rho = -0.65, two-sided p = 0.004), which does not measure adoption, deployment, maturity, or clinical use. First-author geography described publication metadata and discourse, not national attitudes or research quality. Reviews were the least critical and most favourable format. In exploratory forward validation, 2 of 78 early Advocacy predictions were fully borne out, although the analysis was single-rater and retrieval-dependent. Conclusions: Published medical-AI abstracts became modestly less promotional and more focused on specific errors and safety concerns. Unqualified promotion declined, but qualified favourable framing remained dominant, and the rise in critical stance was modest. Concern moved toward errors and patient safety, with factual error discussed more often than fabrication. These findings describe published discourse, not AI capability or whether the evaluations were correct.
Kane, M.; Greene, E. J.; Esserman, D.; Latham, N. K.; Min, L. C.; Ganz, D. A.
Show abstract
Objective: To develop and validate a supervised text-embedded transformer matching model to identify fall injuries in Medicare data, and evaluate the model's performance -- alongside a validated rule-based algorithm-- against "ground truth" from an external reference standard (self-reported fall injuries leading to medical attention). Materials and Methods: Text embeddings of ICD-10-CM and CPT codes in Medicare claims/encounters from participants in the Strategies to Reduce Injuries and Develop Confidence in Elders (STRIDE) trial served as model inputs. Trained on annotated claims/encounters occurring within +/- one month of self-reported fall injuries leading to medical attention, the transformer model generated a continuous 0-1 probability that each claim/encounter was for a fall injury. The model was then applied to all claims/encounters in STRIDE and compared alongside the rule-based algorithm to the external reference standard. Results: The model achieved an area under the curve (AUC) of > 0.96 against annotated claims/encounters in 9 out of 10 holdout folds and 0.85 in the remaining fold. In the full STRIDE dataset, the model achieved a peak AUC of 0.86 (95% CI, 0.84-0.87) against the external reference standard, with results comparable to the rule-based algorithm. Discussion: Relative to rule-based approaches, which typically generate binary outcomes, the continuous event probability generated by the transformer model could support clinical endpoint adjudication, with high-probability predictions treated as events, moderate-probability predictions being adjudicated, and low-probability predictions treated as non-events. Conclusion: A text-embedded transformer model identified fall injuries with comparable accuracy to a rule-based algorithm, demonstrating "proof of concept" for use in endpoint adjudication.